BMJ Health & Care Informatics
● BMJ
Preprints posted in the last 7 days, ranked by how well they match BMJ Health & Care Informatics's content profile, based on 15 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Chowdhury, A. R.; Chowdhury, B.
Show abstract
Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.
Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.
Show abstract
BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.
Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.
Show abstract
The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Jafree, D. J.; Sun, M.; Stewart, G. W.; Gishen, F.; Swanton, C.; Motallebzadeh, R.; UCL MB-PhD Outcomes Study Group,
Show abstract
Background: Clinician-scientists translate clinical observation into discovery, trials, and policy, yet this workforce is shrinking across health systems worldwide. Integrated MB-PhD training, pausing medical training to complete a PhD before clinical exposure or specialisation, is one route into this career. We aimed to evaluate the long-term value of MB-PhD training and the barriers to clinical-academic careers these face after graduation. Methods: We evaluated all 131 graduates (29.8% female) who entered the University College London (UCL) MB-PhD programme over a 25-year period (1994-2018). Bibliometric outputs were collated via an inter-linked information system. Concurrently, all 131 graduates were invited to respond to open-ended questions on career benefits and structural barriers; 99 (75.6%) responded, and responses were independently coded into themes, which were then reviewed and confirmed by a Study Group of 107 individuals, including the 91 respondents who agreed to participate further. Results: Graduates produced 5,877 publications (1,141 first-author, 819 corresponding-author), attracting 350,754 citations, with a mean relative citation ratio of 3.30 {+/-} 0.47, approximately three times the field average and sustained across three decades of programme entry. Graduates secured an estimated $157.55 million across 99 grants, released 465 public datasets, and were named investigators on 31 clinical trials across five continents. Among the 99 survey respondents, 49.5% held consultant-grade posts, 72.7% remained research-active, and 25.3% had reached senior academic grade. Open-ended responses were coded into five recurring structural barriers, subsequently confirmed by the Study Group: insufficient protected research time (72.2% of responses), unsupportive training structures and limited career opportunities (36.7%, 24.4% of responses), funding and pay barriers (22.2% of responses), and lack of mentorship or geographical/family constraints (14.4%, 13.3% of responses). Conclusions: Integrated MB-PhD training generates sustained academic productivity and leadership, but structural barriers threaten retention of graduates within clinical-academic careers. Protecting research time, stabilising funding and pay, and reducing geographic instability are needed to retain the clinician-scientists that health systems have already invested in training.
Zhang, Z.; Qadir, M. I.; Ramchand, R.; Belwadi, M.; Ball, R. P.; Konstantinopoulos, K.; Abbey, E. M.; Ernsberger, K. T.; Guzman, M. J.; Hendren, S.; Holcomb, B. K.; Robb, B. W.; Stankowski, T.; Waters, J. A.; Stefanidis, D.; Bilimoria, K. Y.; Mohanty, S.; Kolbinger, F. R.
Show abstract
Surgical video interpretation is a promising medical artificial intelligence application. However, no existing video annotation method preserves the spatiotemporal complexity of surgeon reasoning. Here we show that verbal reasoning and visual attention can be converted into structured, machine-actionable records of intraoperative behaviours. Our method decomposes transcribed verbal commentary into video-anchored semantic feedback chunks, which are classified via a large language model, with spatial grounding to surgical scenes via eyegaze or cursor tracking. We demonstrate method validity and scalability on structured and unstructured annotation tasks. For quality feedback on full-length colorectal procedures, the method reached near-human fidelity for chunking (mean cosine similarity: 0.95, SD: 0.01) and semantic classification across observations (mean Cohen's kappa: 0.71, SD: 0.07) and evaluative triggers (mean Cohen's kappa: 0.67, SD: 0.14), with excellent usability ratings. For structured critical view of safety assessment in laparoscopic cholecystectomy, implicit annotation yielded excellent agreement with explicit reviewer ratings (Cohen's kappa: 0.83, 0.49 and 0.81 across three criteria). We anticipate this method will advance surgical data science by enabling scalable construction of meaningfully annotated surgical video datasets.
Gorobets, O.; Vinh-Hung, V.
Show abstract
Background: Prostate cancer enzalutamide treatment is approved at a standard dose of 160 mg daily. Concerns for real-world patients -- older and more fragile than those enrolled in clinical trials -- have prompted consideration of initiating treatment with lower doses, but the long-term efficacy of this approach remains unknown. We evaluate the long-term survival and longevity in patients treated with standard versus upfront low-dose enzalutamide. Methods: Retrospective analysis of 151 patients treated with enzalutamide (102 receiving 160 mg; 49 receiving [≤]80 mg) between 2014--2021 at the Centre Hospitalier Universitaire de Martinique, with complete follow-up through end of life (98.7% completeness of follow-up). Primary outcomes were overall survival (OS), progression-free survival (PFS), and longevity (attained age). Results: Doses [≤]80 mg were associated with longer median OS (36.3 vs. 20.7 months), improved restricted mean OS (difference of 0.7 years, p=0.05), and enhanced longevity (median 82.5 vs. 78.3 years, p=0.004). PSA response rate at 12 weeks was higher with lower-dose (71.4% vs. 48.8%, p=0.016). In multivariable models adjusted for prognostic factors, [≤]40 mg compared with 160 mg was non-inferior regarding OS (HR=0.61, 95% CI 0.36--1.06), superior regarding PFS (HR=0.59, 95% CI 0.35--0.99), and superior regarding longevity (HR=0.48, 95% CI 0.28--0.84). Bone metastasis, poor performance status, PSA response, time to PSA nadir, and disease duration were independent predictors of outcomes. A post-hoc analysis revealed a strong association between dose and physician-prescribing profiles, ranging from "endorse-lowest-dose" to "never-deviate-from-full-dose". Conclusions: Lower doses of enzalutamide were non-inferior to full-dose. Dose-adapted strategies warrant further investigation.
Jawhara, B.; Baatiema, L.
Show abstract
Background: Cancer is a growing public health challenge in Ghana, with 27,385 new cases and 17,944 deaths recorded in 2022. Ghana developed a National Cancer Control Strategy (NCCS) in 2011 to guide prevention, early detection, treatment, and palliative care. The strategy expired in 2016 and has not been formally evaluated or renewed, leaving cancer control efforts without a guiding policy framework for nearly a decade. This study examined how the strategy was implemented, what barriers were encountered and what stakeholders recommend for a strengthened national cancer response. Methods: We conducted a qualitative descriptive study using semi-structured key informant interviews. Fifteen participants were recruited through purposive sampling, supplemented by snowball referrals, representing three groups: Ministry of Health policymakers, frontline healthcare providers and representatives of cancer-focused non-governmental organisations. Data were collected between June and September 2025 and analysed using Braun and Clarke's six-phase thematic analysis framework, guided deductively by the WHO Health Systems Building Blocks framework Results: Three themes emerged: NCCS interventions and systems implemented, capturing progress in cancer awareness, HPV vaccination and pilot screening programmes alongside persistent geographic and financial inequities in access; barriers to implementation, including inadequate financing, infrastructure and workforce shortages, the absence of a national cancer registry and governance failures, among them the finding that no frontline healthcare provider interviewed had any awareness of the NCCS; and recommended implementation strategies, including co-production of a renewed strategy, establishment of a dedicated National Cancer Control Programme, expanded health insurance coverage and decentralisation of oncology services. Conclusion: The NCCS was not operationally embedded in the health system. The evidence points to failures in policy dissemination as a constraint that precedes resource constraints. Addressing Ghana's rising cancer burden requires renewed political commitment, co-produced governance structures and accountability mechanisms. These findings have relevance for other low- and middle-income country settings facing similar challenges.
Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.
Show abstract
Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.
Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.
Show abstract
Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.
Song, Q.; Ni, C.; Liu, W.; Li, Y.; Malin, B. A.; Yin, Z.
Show abstract
Automatic coding from clinical notes has been studied extensively for International Classification of Diseases (ICD) codes, yet broad Current Procedural Terminology (CPT) and Healthcare Common Procedure Coding System (HCPCS) recommendation remains comparatively underexplored. Existing studies often focus on one specialty, a limited code vocabulary, or a single model family, leaving it unclear how different artificial intelligence (AI) paradigms perform under a common, clinically meaningful evaluation. We formulate CPT and HCPCS coding as an AI-assisted recommendation task in which a physician or professional coder reviews a short, ranked list of candidate codes supported by the clinical note. Using operative notes from Vanderbilt University Medical Center (VUMC) and discharge summaries from Medical Information Mart for Intensive Care IV (MIMIC-IV), we compare lexical retrieval, Clinical-Longformer, GPT-5.6-Sol, MedGemma-27B, and an inspectable agentic-style retrieve-and-verify system under a controlled review budget. Micro-averaged recall within a fixed number of recommendations measures whether reference codes reach the reviewable list; micro-F1 is reported only where reference labels are sufficiently complete. Zero-shot GPT-5.6-Sol achieves the highest recall within five and ten candidates: 0.717 and 0.800 on VUMC and lower-bound values of 0.689 and 0.738 on MIMIC-IV. The retrieve-and-verify system reaches 0.695 and 0.784 on VUMC and lower-bound values of 0.575 and 0.657 on MIMIC-IV, with a candidate-linked evidence window attached to each retained recommendation. Diagnostic analyses reveal distinct failure sources, including output-length underfilling, confusion among closely related codes, out-of-knowledge-base generation, and incomplete evidence support. These findings establish a systematic evaluation framework for procedure-code recommendation and identify practical requirements for future systems that are accurate, review-efficient, and grounded in clinical evidence.
McHenry, R. D.; Moultrie, C. E.
Show abstract
Objectives Emergency Department (ED) crowding is an international concern, predominantly caused by 'exit block', the lack of availability of inpatient beds for those requiring admission. The implementation of Flow Navigation Centre Plus (FNC+) services in Scotland aimed to reduce self-presentation to EDs and reduce crowding by providing remote clinical assessment for patients contacting urgent care by telephone and professional-to-professional advice on patient pathways, but their effectiveness is unknown. This study aimed to estimate the effect of board-wide implementation of FNC+ on ED attendances and long waits during the first year of FNC+ operation. Methods Controlled interrupted time series using weekly, publicly reported Public Health Scotland data. The intervention was implementation of the FNC+ in NHS Lanarkshire on 1 April 2024. Counts were summed across constituent sites and percentages derived from board totals. Co-primary outcomes were ED attendance volume and the proportions of attendances spending more than 4, 8 and 12 hours in the department. Segmented regression was fitted with contemporaneous control boards, seasonal terms, and accounted for autoregression. Results 118 pre-intervention and 52 post-intervention weeks were analysed across all 3 EDs in the implementing board. Attendances showed no detectable step change (+1.20%; 95%CIs -0.66 to +3.10) relative to the counterfactual. The estimated effect increased across follow-up, however, changing by +3.95% over 52 weeks (95% CI +0.36 to +7.67%). There was no significant step change in the proportion of attendances waiting more than 4 hours following the intervention (+1.74%; 95%CIs -0.71 to 4.20%). Some transition and structural sensitivity analyses demonstrated significant deteriorations in ED performance, and increased attendances, in the year following implementation, and none demonstrated improvements. Conclusions Board-wide implementation of a Flow Navigation Centre Plus was not associated with a step change in ED attendances or in long waits, but there is some evidence that attendances increased and long waits increased in the year following implementation. Their provision of supply-sensitive care is a possible mechanism. Additionally, given their action at the point of input, aiming to divert patients from ED attendance, it is unlikely that such services could relieve a constraint due to exit block, the availability of inpatient care for those requiring admission.
Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.
Show abstract
Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.
Perlman, A.; Goldstein, N.; Goldman, M.; Shapiro, M.; Barash, E.; Bar, A.; Raveh, T.; Tordjman, E.; Schussheim, H.; Dormont, F.; Matalon, O.
Show abstract
Background. Cardiovascular-outcomes trials are lengthy, costly, and associated with substantial uncertainty prior to readout. In-silico trial simulation using real-world data (RWD) has emerged as a potential tool to support earlier decision-making; however, evidence of prospective predictive validity, generated prior to trial result disclosure, remains limited. Methods. We applied a semi-mechanistic machine learning framework integrating real-world patient data with biologically informed drug representations to prospectively simulate the VESALIUS-CV trial evaluating evolocumab versus placebo. The simulation model was trained on a combination of patient-level real-world data and a drug-centric knowledge graph and validated for both patient-level and trial-level retrospective predictive performance. The model was then used to simulate VESALIUS-CV before public disclosure of trial results, using a locked model and prespecified eligibility criteria and primary endpoint aligned with the clinical protocol. A patient-level time-to-event model was used to generate virtual trial arms, from which cumulative incidence curves, hazard ratios, confidence intervals, and p-values for major adverse cardiovascular events (MACE) were estimated. Results. In retrospective validation, the model demonstrated strong patient-level discrimination, with time-dependent ROC-AUC values ranging from 0.80 to 0.90 across follow-up horizons. For trial-level validation, 22 randomized cardiovascular-outcomes trials were simulated, and hazard ratios for 3-point MACE across 24 between-arm comparisons showed consistent directional agreement and quantitative correlation with published results such that the model accurately predicted trial success, achieving an F1 score of 0.83, with precision of 0.79 and sensitivity of 0.89. In a fully prospective application, the simulation predicted a statistically significant reduction in 3-point MACE with evolocumab versus placebo, estimating a hazard ratio of 0.78 (95% CI, 0.70-0.87) at 54 months. These predictions were consistent with the subsequently reported VESALIUS-CV results, which demonstrated a hazard ratio of 0.75 (95% CI, 0.65-0.86) at 55 months of median follow-up. Conclusions. In a fully prospective setting, a RWD-driven, AI-based simulation accurately predicted the direction, magnitude, and temporal dynamics of treatment effects observed in the VESALIUS-CV trial. These results demonstrate that in-silico trial simulation can anticipate clinical outcomes in the prospective setting, supporting its use as a complementary tool for early decision-making, trial design optimization, and de-risking in cardiovascular drug development.
Mathew, Z.; Mehta, R.; Kim, S.; Jeyaraj, J.; Asif, T.
Show abstract
Background: Primary malignant cardiac tumors (PMCTs) are rare and histologically heterogeneous. Objective: To compare demographics, specific ICD-O-3 morphologies, first-course treatment patterns, annual registered case counts, and unadjusted overall survival between soft-tissue and hematologic PMCTs. Methods: We identified 730 PMCT cases diagnosed from 2000 to 2021 in SEER 18 (ICD-O-3 topography C38.0). Histologic lineage was assigned from ICD-O-3 morphology. Comparative analyses included soft-tissue (n=458) and hematologic (n=212) tumors. First-course variables were primary-site surgery, chemotherapy (yes versus no/unknown), and radiotherapy (radiation versus none/unknown). Groups were compared with chi-square tests. Overall survival was estimated with Kaplan-Meier methods; follow-up was truncated at 120 months. Results: Soft-tissue PMCTs occurred predominantly at ages 45-64 years (67.9%), whereas hematologic PMCTs occurred predominantly at age [≥]65 years (63.2%; p<0.001). Men comprised 59.9% of hematologic and 49.3% of soft-tissue cases (p=0.014). The leading soft-tissue morphology was hemangiosarcoma/angiosarcoma (ICD-O-3 9120/3; 201/458, 43.9%); synovial sarcoma accounted for 20/458 cases (4.4%). Diffuse large B-cell lymphoma, NOS, accounted for 131/212 hematologic tumors (61.8%). Any primary-site surgery was recorded in 66.6% of soft-tissue versus 15.6% of hematologic cases (p<0.001). Chemotherapy was recorded in 67.5% versus 51.1% (p<0.001), and radiotherapy in 9.0% versus 20.5% (p<0.001). In exploratory Kaplan-Meier analyses, hematologic patients with recorded chemotherapy had higher unadjusted 120-month overall survival than those without recorded chemotherapy (42.0% versus 12.2%; log-rank p=7.5x10-). Radiation-associated survival differences were not statistically significant in either lineage. Conclusions: Soft-tissue and hematologic PMCTs have distinct age distributions, named histologies, and first-course treatment patterns in SEER. These findings describe registry coding and do not establish treatment effectiveness or population incidence.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Witham, M.; Evison, F.; Bellass, S.; Cooper, R.; Gallier, S.; Pretorius, S.; Sapey, E.; Suklan, J.; Sayer, A. A.
Show abstract
Study Objective Little is known about where in hospital care for multiple long-term conditions (MLTC) is delivered. We aimed to describe pathways of care (ward transfers) and outcomes for people admitted to hospital for unscheduled care by MLTC status and other key sociodemographic characteristics. Design and setting Analysis of routinely-collected electronic health records from a large acute UK hospital. Participants Adult unscheduled care admissions from 1st July 2018 to 30th June 2019. The presence of two or more of 59 long-term conditions was ascertained using ICD-10 codes from previous hospital discharges. Main outcome measures Markov state transition probabilities were derived for ward moves and compared for MLTC vs no MLTC, age, sex, ethnicity and neighbourhood deprivation. Outcomes (length of stay, death, readmission, move from definitive ward) and time spent in emergency and assessment departments were compared between subgroups. Results A total of 33,252 adults, mean age 56.0 (SD 21.9) years were analysed; 14,834 (42.4%) had MLTC. People with MLTC were more likely to die in hospital (4.2 vs 1.9%, p<0.001), transfer to internal medicine wards or older peoples medicine wards, were less likely to transfer to surgical wards, had longer median length of stay (1.83 vs 0.69 days, p<0.001), stayed longer in acute medical units (15.5 vs 9.6 hours, p<0.001), and were more likely to move from their definitive ward (18.2 vs 16.4%, p=0.002). Conclusion Unscheduled hospital care pathways are complex and differ for people with MLTC, who have worse outcomes and may be less likely to receive optimal care.
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
Yano, Y.; Shintani, E.; Arita, S.; Ashine, R.; Iinuma, N.; Mori, H.; Fujibayashi, K.; Yamada, Y.; Saita, M.; Nakashima, N.; Itoh, H.; Nangaku, M.; Ohashi, M.; Daida, H.; Arai, H.; Naito, T.
Show abstract
The widespread adoption of clinical large language models (LLMs) introduces significant risks of automation bias, premature closure, and clinician deskilling. Current interpretability paradigms, including latent space trajectories, Concept Activation Vectors, and Concept Bottleneck Models, suffer from topological stagnation, metric distortion, and epistemic occlusion, frequently masking intermediate diagnostic uncertainty behind falsely confident outputs. To address these structural vulnerabilities, this paper introduces a novel closed-loop, multi-agent framework designed to quantify and visualize dynamic epistemic uncertainty in clinical LLM reasoning. By coupling predictive Shannon entropy with non-linear Isometric Feature Mapping (ISOMAP), the architecture projects high-dimensional inference state vectors onto a calibrated two-dimensional latent space, thereby assigning a quantifiable thermodynamic energy state to the reasoning path to track diagnostic velocity, cognitive momentum, and trajectory efficiency across sequential diagnostic rounds. Pilot validation across representative emergency medicine scenarios demonstrated distinct topological and information-theoretic behaviors: unconfounded cases (cerebellar infarction) exhibited smooth geodesic progression toward the ground truth alongside monotonic Shannon entropy decay from 2.15 to 1.74; noisy environments with ambiguous findings (spontaneous pneumothorax) suffered from trajectory wandering, local minimum traps, and high sustained entropy (~2.41) due to insufficient repulsive weighting for negative evidence; and triage-conflicted cases (acute cholangitis) achieved precise geometric proximity to the true node but experienced top-1 rank stagnation because the model conflated acute severity triage (sepsis) with anatomical etiology. By rendering machine hesitation and cognitive divergence visually auditable before final diagnostic crystallization, this geometric-information framework enables dynamic trust calibration and human-AI co-regulation at the point of care while establishing a clear mathematical foundation for future architectural interventions, such as dual-channel safety decoupling and non-linear repulsive weighting. Moving forward, validating these architectural enhancements across large-scale electronic health record databases and prospective clinical trials will be essential to realize its full clinical utility, establishing a foundational blueprint for safe, transparent, and cognitively synergistic AI integration in future medical practice. By rendering the LLM's reasoning process visually auditable, this framework lays the groundwork for capturing and externalizing the clinician's own cognitive patterns within the AI, forming a coupled system. This enables the explicit visualization of cognitive gaps between physician hypotheses and AI inferences, transforming the interaction from simple answer-checking into a dynamic learning process for both human and machine that prevents diagnostic oversight. Ultimately, because the responsibility for final clinical decision-making remains with the human practitioner, this framework serves as a vital decision-support mechanism. Moving forward, validating these architectural enhancements across large-scale electronic health record databases and prospective clinical trials will be essential to realize its full clinical utility, establishing a foundational blueprint for safe, transparent, and cognitively synergistic AI integration in future medical practice.